Back

Protein Science

Wiley

Preprints posted in the last 30 days, ranked by how well they match Protein Science's content profile, based on 246 papers previously published here. The average preprint has a 0.17% match score for this journal, so anything above that is already an above-average fit.

1
ResiRuler: A Toolkit for Visualizing Residue-Residue Distances and Structural Changes in Biomolecular Models

Baker, T. H.; Ohi, M. D.; Salmen, W.

2026-08-23 bioinformatics 10.64898/2026.08.19.745761 medRxiv
Top 0.1%
26.4%
Show abstract

Proteins and their associated complexes often adopt multiple conformations, with the transitions between these states playing a critical role in biological function. However, the resulting structural heterogeneity can be challenging to visualize and communicate, often requiring manual inspection and time-consuming annotation of biomolecular structures. To address this, we developed ResiRuler, a local, browser-based tool that uses inter-residue distance measurements to quickly quantify atomic displacements and map changes in internal geometry across ensembles of related protein structures. By converting structural differences into residue-pair distance changes, ResiRuler enables rapid identification of regions undergoing coordinated motion, local rearrangement, or large-scale conformational change. The resulting visualizations can be exported as scripts for PyMOL and ChimeraX, allowing users to explore conformational differences and generate publication-quality molecular figures in their preferred visualization environment. Using atomic models in Macromolecular Crystallographic Information File (mmCIF) file format, ResiRuler aligns multiple structures and measures structural variation across models facilitating visualization and presentation of these differences. This allows for rapid visualization of which regions of proteins change among ensembles of structures. The program is available for download at https://github.com/tbaker67/ResiRuler on macOS and Linux operating systems.

2
Large-scale structure prediction of DUF-containing protein-protein interactions

Riepenhausen, L.; Costa, F.; Andreeva, A.; Bateman, A.

2026-08-20 bioinformatics 10.64898/2026.08.19.745780 medRxiv
Top 0.1%
19.1%
Show abstract

Motivation: Continuing advances in genome and metagenome sequencing expand the number of identified conserved protein families that remain functionally uncharacterized and contain domains of unknown function (DUFs). Functional-association resources such as STRING provide biological context, but mostly do not distinguish indirect association from physical interaction. We assessed whether AlphaFold 3 complex prediction, combined with STRING evidence and domain-level analysis of interfaces and interaction partners, can help identify and characterize DUF-containing proteins. Results: We generated four structural-prediction cohorts from STRING associations involving DUF-containing proteins and evaluated the predicted complexes using interface ipSAE, average pLDDT and buried surface area. An L2-regularized logistic regression model was trained on an initial cohort of predictions from high-confidence STRING associations to prioritize DUF-containing candidates likely to produce structurally confident AlphaFold 3 complexes. The model was then applied across all 12,535 organisms represented in STRING v12.0, followed by grouping into DUF-family and partner-architecture modules, covering 2,076 unique DUF families. The final L2-model screen contained 12,298 successfully modelled protein pairs, including 1,208 (9.82%) complexes meeting a strict-confidence criterion and 2,433 (19.78%) meeting a more liberal confidence criterion. Two examples suggest roles for DUF4130 in nucleic-acid-associated radical-SAM biology and DUF5819 in a bacterial system related to vitamin-K-dependent carboxylation. Availability and implementation: Predicted structures and associated metadata are available through Zenodo at https://doi.org/10.5281/zenodo.21875362. The model implementation and code used to generate the analyses and figures are available at https://github.com/linoriep/Proteome-scale-structure-prediction-of-DUF-containing-protein-protein-interactions.

3
RheoScale 2.0: Revealing the Hidden Roles of Protein Positions via Substitution Patterns

Liu, D.; Sreenivasan, S.; Gray, C. J.; Cleveland, H. C.; Swint-Kruse, L.

2026-08-11 biochemistry 10.64898/2026.08.10.743964 medRxiv
Top 0.1%
18.8%
Show abstract

A central challenge in molecular biology is understanding how amino acid substitutions modulate various features of protein function and stability. To illuminate the complexities of this relationship, high-throughput (HTP) assays are increasingly used to assess site-saturating mutagenesis libraries. A common downstream analysis is to average the set of twenty outcomes at each amino acid position for comparison with structural and evolutionary features. Average values clearly identify positions that tolerate most substitutions (neutral positions) and positions where most substitutions abolish activity (toggle positions). However, average values conceal the existence of rheostat positions, where different amino acid substitutions sample a wide range of outcomes. To quantitatively identify rheostat positions, we previously developed a histogram-based analysis that we here expand by: (i) incorporating new position classes observed in experimental studies of rheostat positions; (ii) formalizing a hierarchy of class assignments; (iii) refining error-based identification of neutral positions; and (iv) statistically assessing the robustness of class assignments to changes in experimental and computational parameters. RheoScale 2.0 is implemented in Excel and newly implemented in Python for facile integration with existing HTP pipelines; all parameters are customizable. Example analyses are shown for three HTP datasets of the SARS-CoV-2 papain-like protease. Results illustrate two aspects that influence interpretation of HTP data: First, position assignments (and substitution outcomes) depend highly on the measured feature. Second, many protein positions play multiple roles in the sequence-structure-function relationship. The recognition of varied position roles will advance understanding of pathogen evolution, protein engineering, and variant interpretation for personalized medicine. SummaryRheoScale 2.0 improves how high-throughput mutational data are interpreted by identifying protein positions where amino acid substitutions act like biological dimmer switches. By enabling more nuanced assignment of position behavior, beyond neutral or deleterious outcomes, this analysis framework advances studies of sequence-structure-function relationships and has broad relevance for understanding protein evolution, engineering proteins with desired properties, and interpreting variants linked to human disease. SOFTWARE AVAILABILITYhttps://github.com/liskinsk/RheoScale-calculator

4
A thermodynamic framework for mapping elastic recoil mechanism across the human proteome

Desai, R.; Pople, D.; Musale, A.; Jain, S.; Sajjad, I.; Wittebort, R. J.; Koder, R. L.; Nanda, V.

2026-08-30 biophysics 10.64898/2026.08.28.747957 medRxiv
Top 0.1%
15.1%
Show abstract

The folding thermodynamics of proteins are dominated by two opposing forces, the loss in backbone entropy and the packing of hydrophobic groups. The same forces are major contributors to the extension thermodynamics of elastic proteins with the distinction that both processes act in concert, favoring the higher chain and solvent entropy of a relaxed conformation. The relative entropic contributions specify the recoil mechanism; human elastin recoil is primarily driven by hydrophobic forces, whereas fly resilin has a rubber-like mechanism driven by backbone entropy. Despite the importance of elastic proteins to tissue biomechanics, few have been identified, let alone characterized to the same extent as elastin and resilin. We develop a thermodynamic framework that maps proteins by sequence-derived estimates of extension-induced backbone and solvent entropy changes. Putative elastic proteins are proposed and classified by recoil mechanism based on estimated thermodynamic features. Proteins that map to elastic regions are overrepresented by the skin proteome. The set of predicted elastic domains is further extended by incorporating sequence context embedded in protein language models. Protein domains with distinct thermodynamic recoil mechanisms cluster on the latent space manifold. Some of these domains are anticipated to have roles within molecular machines, expanding the scope of elastic protein function beyond mechanical materials like elastin and resilin.

5
Amyloid Polymorphism of Lysozyme Governs Cross-Seeding of Insulin Aggregation

Metkar, S.; Eerati, V.; Ramamoorthy, A.

2026-08-30 biophysics 10.64898/2026.08.26.747312 medRxiv
Top 0.1%
14.9%
Show abstract

Amyloid fibrils are highly ordered protein aggregates characterized by a conserved cross-{beta}-sheet architecture despite originating from structurally diverse precursor proteins. Growing evidence suggests that interactions between different amyloidogenic proteins can modulate aggregation pathways through heterologous cross-seeding; however, the influence of seed polymorphism on the structure and biological properties of cross-seeded fibrils remains poorly understood. Here, we investigated the cross-seeding of native human insulin by two structurally distinct polymorphs of hen egg-white lysozyme (HEWL): flexible fibrils (FFs) and rigid fibrils (RFs). Native insulin remained stable under physiological conditions and underwent spontaneous fibrillation only under acidic conditions. In contrast, both HEWL polymorphs efficiently induced insulin aggregation at physiological pH, bypassing the nucleation barrier. Thioflavin T fluorescence, circular dichroism spectroscopy, and transmission electron microscopy revealed that lysozyme FFs templated the formation of insulin flexible fibrils (IFFs), whereas lysozyme RFs produced insulin rigid fibrils (IRFs), demonstrating that the structural characteristics of the parental HEWL polymorphs were propagated during heterologous cross-seeding. The toxicity of the resulting insulin fibrils was evaluated in SH-SY5Y neuronal cells and CCF-STTG1 astrocytes. IFFs exhibited minimal cytotoxicity and only subtle morphological alterations, whereas IRFs caused modest reductions in cell viability accompanied by more pronounced cellular damage. These findings demonstrate that the structural polymorphism of HEWL fibrils governs both the architecture and biological activity of cross-seeded insulin fibrils, highlighting amyloid polymorphism as an important determinant of heterologous amyloid propagation and a potential design principle for engineering functional amyloid-based biomaterials and protein delivery platforms.

6
Makeshift: a lightweight software for accessing and analyzing NMR data and protein dynamics

El Nesr, G.; Wayment-Steele, H. K.

2026-08-20 biophysics 10.64898/2026.08.17.745346 medRxiv
Top 0.1%
14.8%
Show abstract

Nuclear magnetic resonance (NMR) spectroscopy yields rich residue-level information on biomolecular dynamics and chemical environments, two frontiers for quantitative predictive methods in biochemistry. Decades of data are publicly archived in the Biological Magnetic Resonance Data Bank (BMRB)1, yet in practice, this information remains difficult to access and interpret at scale and within computational workflows. Here we present makeshift, an open-source Python package for accessing, curating, and analyzing NMR datasets. Users can readily retrieve and parse BMRB entries and perform essential analyses such as chemical shift re-referencing, secondary structure propensity prediction, and interpretation of relaxation datasets for dynamics. We re-implemented several widely-used NMR data calculations which were not open-source or available in Python and validated our implementations against the original implementations. By integrating data access, processing, and analysis into a single Python interface, makeshift lowers the barrier for reproducible, scalable analysis and machine learning applications using biomolecular NMR data.

7
Hydration Energetics Shape Antibody Discrimination between Sulfotyrosine and Phosphotyrosine

Mori, T.; Yahagi, K.; Maruoka, S.; Toyoda, K.; Sonoshita, Y.; Kametani, Y.; Shiota, Y.; Yoshizawa, K.; Watanabe, K.; Okazaki, K.; Kobashigawa, Y.; Morioka, H.; Hirakawa, H.; Nishimoto, E.; Teramoto, T.; Kakuta, Y.

2026-08-11 biophysics 10.64898/2026.08.05.743142 medRxiv
Top 0.2%
12.8%
Show abstract

Chemically similar post-translational modifications can mediate distinct biological functions, but how proteins distinguish between them remains unclear. Sulfotyrosine (sTyr) and phosphotyrosine (pTyr) exemplify this problem because they have similar sizes, local geometries, and electrostatic properties but function in different biological contexts. Here, we used the monoclonal antibody PSG2, which recognizes sTyr independently of the surrounding peptide sequence, to examine how a protein distinguishes these modifications. The crystal structure of PSG2 bound to an sTyr-containing peptide revealed a deep electropositive pocket with no modeled water molecules in direct contact with the sulfate group. Gas-phase density functional theory calculations favored pTyr over sTyr, showing that direct protein-ligand interactions alone are insufficient to explain PSG2 selectivity. Explicit first-shell hydration calculations showed that pTyr has a larger desolvation penalty than sTyr, and accounting for this difference reversed the calculated energetic order. Isothermal titration calorimetry showed favorable enthalpic and entropic contributions to sTyr binding, whereas no detectable heat signal was observed for pTyr. These results show that PSG2 distinguishes sTyr from pTyr through the balance between direct protein-ligand interactions and ligand desolvation.

8
PyMOL plugin for Protein Circuit Topology

Dimins, M.; Bazba, A.; Mogyorosi, A.; Kennon, E.; Fiol, T. D.; Hagen, L. A.; Sheikhhassani, V.; Akulov, V.; Mashaghi, A.

2026-08-20 bioinformatics 10.1101/2025.10.21.683762 medRxiv
Top 0.2%
12.8%
Show abstract

Circuit Topology (CT) provides a fundamental framework for analysing folded polymer chains, with applications in functional annotation, protein engineering and drug development. We present a protein CT analysis plugin for PyMOL v3.1.6.1 with a graphical user interface (GUI), automatic installation, and novel features developed through integration with PyMOL's application programming interface (API). The plugin integrates various previously developed CT methodologies for studying structured proteins and their complexes as well as the dynamics of disordered proteins. Analysis of a representative protein and a molecular dynamics trajectory demonstrates the plugin's three analysis modes and their outputs. The plugin reproduces the reference ProteinCT implementation exactly on the structures tested, and is distributed with a versioned release, a pinned environment and a one-command reproduction of every result reported here.

9
NMR assignments and secondary structure analysis of the human 5MP1 C-terminal domain

Seker, A.; Anand, S.; Marintchev, A.

2026-08-18 biophysics 10.64898/2026.08.11.744028 medRxiv
Top 0.2%
12.5%
Show abstract

Eukaryotic translation initiation is tightly regulated by interactions among translation initiation factors (eIFs) that ensure accurate start codon selection. The translation regulator, eIF5 mimic protein 1 (5MP1) contributes to this process by competing with eIF5 for binding to eIF2, thereby increasing the stringency of translation initiation. Despite its important regulatory role and emerging involvement in tumorigenesis, structural information on human 5MP1 remains limited. Here, we report the near-complete backbone and partial side-chain NMR resonance assignments of the C-terminal domain of human 5MP1 (residues 250-419), carrying a W404E substitution that disrupts dimerization. The WT protein forms a dimer at NMR concentrations, which increases the effective size of the protein and also causes disappearance of peaks corresponding to aminoacids at the dimer interface due to conformational exchange. Backbone resonance assignments were completed for 96.4% of the non-proline residues. Secondary structure was analyzed using Chemical Shift Index (CSI) and compared with the AlphaFold structural model. Regions of disagreement between the experimental and computational secondary structure assignments were further examined using 15N-NOESY-HSQC spectra, allowing experimental validation of local structural features. While the AlphaFold model accurately reproduces the overall fold of the 5MP1 C-terminal domain, several localized discrepancies were identified, particularly near the N- and C-terminal regions of the domain, where experimental NMR data support alternative secondary structure assignments. These resonance assignments and experimentally validated structural features provide a foundation for future investigations of the molecular interactions, dynamics, and functions of 5MP1 in translation initiation.

10
Computational Structural Analysis of POLG Variants R627Q and W748S with Model-Variability Controls

Friedl, A.; Manst, D.

2026-08-27 biophysics 10.64898/2026.08.25.747108 medRxiv
Top 0.2%
11.9%
Show abstract

Background: Comparisons between independently predicted wild-type and missense-variant protein structures can generate mechanistic hypotheses, but small apparent differences may reflect model-selection variability rather than mutation-specific effects. Methods: Human mitochondrial DNA polymerase gamma (POLG; UniProt P54098) variants p.Arg627Gln (R627Q) and p.Trp748Ser (W748S) were evaluated using five AlphaFold2-PTM network-model outputs per condition generated with one random seed under matched ColabFold settings. Ten pairwise wild type comparisons at each site described between-network model-selection variability. Variant effects were summarized across five within-network wild-type-versus-variant comparisons using rotation-invariant local C-alpha pair distances and local displacement after global and local alignment. Because these comparison designs differ, the wild-type distribution was used as context rather than a mutation-effect null. Wild-type cryo-EM structure 9GGF was used for contact and interface mapping. Experimental A467T and G848S structures 9GGE and 9GGC provided contextual benchmarks. Results: R627Q measurements fell within the range of between-network wild-type differences: its median mean local pair-distance change was 0.170 angstrom, compared with a wild-type median of 0.170 angstrom, and its locally aligned displacement was 0.265 versus 0.248 angstrom. W748S showed higher median values (0.168 versus 0.132 angstrom for pair-distance change; 0.236 versus 0.182 angstrom for locally aligned displacement), but the ranges overlapped and the comparison-design asymmetry precluded a calibrated mutation-effect percentile. Experimental A467T and G848S comparisons produced local changes of similar magnitude. In 9GGF, R627 and W748 directly shared a local microenvironment, with a minimum heavy-atom distance of 3.53 angstrom. R627 also formed short polar-contact candidates with D629 and D743, whereas W748 occupied a hydrophobic packing environment containing Y622 and F750. Both sites were more than 18 angstrom from nucleic acid, more than 30 angstrom from POLG2, and more than 33 angstrom from PZL-A in a ligand-bound structure. Conclusions: Available AlphaFold2 comparisons do not establish a mutation-specific structural deformation for either variant. Experimental-structure mapping supports testable physicochemical hypotheses involving a shared R627-W748 microenvironment - loss of an arginine-centered polar network for R627Q and disruption of a buried aromatic environment for W748S - but not direct DNA, POLG2, or PZL-A contact mechanisms. Matched control substitutions and independent seeds are required to calibrate small mutation-associated structural deltas.

11
Molecular dynamics descriptors for 1,079 post-translationally modified protein systems

Liu, K.; Qian, Q.; Peng, J.; Ma, D.; Yao, Y.; Zhao, J.; Chi, Y.

2026-08-24 biophysics 10.64898/2026.08.23.746532 medRxiv
Top 0.2%
11.9%
Show abstract

Databases of post-translational modifications (PTMs) catalogue modified sites and increasingly add static structural context, but trajectory-derived descriptors remain scattered across specialised tools and general molecular dynamics archives. Dyna-MO PTM brings together 1,079 AlphaFold 3-seeded systems covering lysine acetylation, lysine and arginine monomethylation, and serine, threonine and tyrosine phosphorylation. Each system is linked to three completed 10 ns replicas generated with CHARMM36m and TIP3P, for 32.37 s of aggregate sampling. A 118-column table joins simulation and quality-control provenance with global relaxation measures, site solvent exposure, rotamers, secondary structure and ionic-contact proxies. Versioned identifiers connect the records to starting structures, trajectories, manifests and analysis scripts. Researchers can use the resource to filter PTM contexts, reproduce descriptors, prioritise longer simulations and evaluate trajectory-analysis or generative methods. The trajectories describe finite-window relaxation rather than equilibrium free energies, kinetics or matched PTM effects.

12
Prot2Surf: fast analysis of protein - surface binding modes

Muniz-Chicharro, A.; Tanriver, G.; Gora, A.

2026-08-29 bioinformatics 10.64898/2026.08.26.747352 medRxiv
Top 0.3%
11.2%
Show abstract

Summary: Prot2Surf is a software tool designed for the characterization and prediction of protein association to surfaces. In this application note, Prot2Surf was tested using catalytic domains of the lytic polysaccharide monooxygenases (LPMOs), interacting with native surfaces. The results show that the software can efficiently analyze key binding features, including protein-surface distances, distances between catalytically reactive atoms, and the orientation angle between surface chains and the protein. These features are essential for distinguishing productive binding poses in these protein-surface systems and for understanding interaction patterns that provide guidance on protein engineering. Prot2Surf performs these analyses within seconds to a few minutes, providing a fast and accessible framework to post-process and characterize protein-surface encounter complexes. Availability and implementation: Prot2Surf, which is written in Fortran90, is documented and freely available as open source on GitHub: https://github.com/TUNNELING-GROUP/Prot2Surf. In order to run Prot2Surf, users should also install the SDA software package which is freely available at https://www.h-its.org/downloads/sda7/.

13
Incidental conformational switching in an allosteric enzyme

Sapienza, P. J.; Vera-Rodriguez, D. J.; Mileur, T. R.; Lee, A. L.

2026-08-07 biophysics 10.64898/2026.08.04.741757 medRxiv
Top 0.3%
10.9%
Show abstract

The classical understanding of allostery was initially grounded in two-state models, such as MWC and KNF, where structure and function are inextricably linked through transitions between low-(T) and high-affinity (R) states. Here, we show Yeast chorismate mutase (CM) provides a vivid example of the growing list of exceptions to the traditional T vs R two-state allosteric paradigm. While CM exhibits dynamic sampling of the R-state in the presence of the activator tryptophan (Trp), suggesting a conformational selection (CS) mechanism, we present multiple instances where conformational status and catalytic activity are decoupled. Using NMR spectroscopy and kinetic assays, we identify CM variants that reside almost exclusively in the T conformation can exhibit maximal activity, while others that predominantly occupy the R conformation are weakly active. Quantitative comparison of experimental data with a parameterized CS model reveals deviations of up to two orders of magnitude, ruling out the simplest two-state model for substrate affinity modulation in this system. We propose that the observed T-to-R switching in CM is "incidental", a byproduct of an evolved energy landscape that allows access to the substrate-bound pose but does not mechanistically determine affinity. Our findings suggest that allosteric regulation in CM may instead be driven by local features of the ground-state ensemble, which operate independently of global T/R status. This work further highlights an emerging view that the mere observation of a pre-sampled active conformation does not sufficiently prove a two-state mechanism and further underscores the need for deeper ensemble-based perspectives in protein engineering and allostery.

14
PepXPro: a framework for curating, generating, and optimizing structure-affinity protein-peptide datasets

Chi, L. A.; Ytreberg, F. M.

2026-08-18 biophysics 10.64898/2026.08.09.743757 medRxiv
Top 0.3%
9.7%
Show abstract

Protein-peptide interactions are central to cellular signaling and to a growing class of peptide therapeutics, yet the datasets used to develop and benchmark computational methods for protein-peptide modeling remain poorly standardized. Available databases prioritize comprehensive coverage but require task-specific curation, while published benchmarks are typically distributed as static collections built with heterogeneous curation, quality-filtering, redundancy-reduction, and sampling strategies, limiting reproducibility and cross-study comparison. We present PepXPro, a modular framework that transforms publicly available protein-peptide structure-affinity resources into curated datasets and reproducible benchmark collections generated under user-defined criteria. PepXPro is organized into three components: Scrape, for deterministic curation of protein-peptide complex entries from public resources; GenSample, for constructing configurable subsets under explicit quality, redundancy, and sampling constraints; and Benchmark, for evaluating candidate subsets and selecting a nonredundant, representative, general-purpose benchmark for distribution. Starting from PDBbind and complementary resources, the curation pipeline yields a pool of proteinpeptide complex entries that retains chemically complex cases, including disulfide- linked cyclic peptides, which are commonly excluded from existing benchmarks. We release PepXPro Benchmark v1, a benchmark comprising 70 non-redundant protein- peptide complexes with experimentally determined structures and binding affinities. The underlying framework provides an extensible foundation for reproducible protein- peptide benchmark construction.

15
Morpheus-3D: Structural Diversity-Guided Detection and Localization of Protein Fold Switching

Kuniyil, S.; Subramanian, V.; Arun, A.; Lakshmanan, A.; Sekhar, A.; Srivastava, A.

2026-08-21 biophysics 10.64898/2026.08.16.745091 medRxiv
Top 0.3%
9.6%
Show abstract

Proteins that reversibly adopt multiple stable folds challenge the classical sequence-structure paradigm, yet their discovery remains limited because fold switching is difficult to detect experimentally and current computational methods fail to resolve the underlying conformationally plastic regions. Here we present Morpheus-3D, a sequence-based framework that quantifies residue-level tertiary structural diversity using entropy profiles derived from the Foldseek 3Di structural alphabet. By capturing variation in tertiary interaction environments rather than secondary structure alone, Morpheus-3D identifies fold-switching proteins while simultaneously localizing the sequence regions responsible for structural transitions. The framework outperforms existing predictors, accurately recovers experimentally characterized switching regions, generalizes to recently discovered natural and engineered fold-switching proteins absent from training, and detects conformational plasticity inaccessible to secondary-structure-based approaches. Application to 57 representative proteomes reveals that fold-switching potential is widespread but enriched in regulatory, pathogenic, and environmentally adaptive lineages. Integration with ancestral sequence reconstruction further un-covers evolutionary trajectories through which conformational plasticity emerges. To make these predictions directly accessible, we implemented Morpheus-3D as an interactive web platform (https://morpheus.slicearrow.com/), in which per-residue entropy profiles, sequence and three-dimensional structure are displayed together and respond as one, allowing predicted fold-switching regions to be mapped onto the structure and exported for downstream analysis. Morpheus-3D provides a scalable framework for discovering metamorphic proteins and investigating the origins, mechanisms, and evolution of structural plasticity directly from sequence.

16
Constrained Generative Design Frameworks For Computational Discovery of Target-Specific DARPin Candidates

Pourbaghi, M.; Elemento, O.; Bradbury, M. S.

2026-08-13 bioinformatics 10.64898/2026.08.07.743551 medRxiv
Top 0.3%
9.6%
Show abstract

Applying unconstrained generative protein models to fixed structural scaffolds can produce systematic design artifacts, including a "Glycine Trap" characterized by the enrichment of glycine at structurally incompatible positions. Furthermore, optimizing sequences against artificial rigid-body docking geometries induces reward-hacking and severe geometric hallucinations. In addition, the highly conserved designed ankyrin repeat protein, or DARPin, scaffold can obscure defects at the engineered binding interface, causing AlphaFold2-Multimer (AF2) to predict nonfunctional protein-target interactions with high confidence. To overcome these limitations, we developed DARPinMPNN, a scaffold-constrained computational pipeline for DARPin candidate discovery. Restricting sequence generation to a validated DARPin design space eliminated these failure modes. A state-aware chimeric multiple sequence alignment strategy was engineered and enabled AlphaFold2-Multimer (AF2) to serve as a high-throughput structural sieve, while AlphaFold 3 (AF3) provided independent structural validation of candidate binders. Using this framework, we identified mesothelin-targeting DARPin candidates with predicted structural confidences (champion ipTM = 0.83) approaching those of a structurally validated picomolar-affinity binder (G3 control, ipTM = 0.89). By revealing extensive discordance between AF2 and AF3 predictions, this work establishes a robust framework for identifying and prioritizing high-confidence DARPin candidates for experimental validation.

17
Ahead of the membrane curve: in silico insights into amyloid-β aggregation

Maximiano, P.; Hashemi, M.

2026-08-25 biophysics 10.64898/2026.08.22.746319 medRxiv
Top 0.3%
9.5%
Show abstract

Membrane surfaces can accelerate amyloid $\beta$ (A$\beta$) aggregation, yet the role of membrane curvature in this process remains poorly understood. Here, we used multi-million atom all-atom molecular dynamics simulations to compare the adsorption, conformational dynamics, and oligomerization of four A$\beta$42 peptides at a planar neuronal membrane and a highly curved lipid vesicle. For both systems, all peptides adsorbed within the first 2 $\mu$s, but their subsequent behavior differed substantially. The curved membrane exhibited a larger area per lipid and more extensive hydrophobic packing defects, allowing A$\beta$42 to penetrate more deeply and form strong contacts with lipid tails through its central hydrophobic core and C-terminal region. These interactions disrupted a solution-formed dimer and limited peptide-peptide association during the simulated interval. Additionally, vesicle-bound peptides adopted more extended conformations with increased $\beta$-structure and $\beta$-hairpin formation compared with peptides at the planar membrane. A$\beta$42 adsorption was also corelated to lipid reorganization in the vesicle. In contrast, the planar membrane supported weaker adsorption and stable dimer-to-trimer growth but showed little large-scale lipid segregation. These findings reveal that curvature reshapes the early A$\beta$42 aggregation landscape by strengthening peptide-lipid interactions, altering aggregation-prone conformations, and reorganizing membrane domains. Membrane geometry should therefore be considered alongside lipid composition in mechanistic models of A$\beta$42 oligomerization and membrane-associated toxicity.

18
Safety First: Input Screening for Protein Design Tools

Palmer, P.; Teran, N.; Wheeler, N.; Yassif, J. M.

2026-08-07 synthetic biology 10.64898/2026.08.04.740855 medRxiv
Top 0.3%
9.4%
Show abstract

As biological AI models become more powerful, practical biosecurity approaches are needed to support beneficial applications while reducing misuse risks. Sequence-similarity-based screening approaches are no longer adequate to safeguard biological AI models because these models can design molecules with novel sequences and structures. Therefore, a screening approach that takes function into account is needed. To address this need, we propose a new screening method for AI-enabled protein binder design tools. Our framework screens protein binding targets, with a focus on the human proteome, as opposed to the binder molecule itself. We constructed a database of 14,541 potentially harmful proteoform targets from the human proteome (7.1% of all human protein proteoforms) classified by biosecurity risk level. To discern structural and functional features, we evaluated constructs with an embedding-based screening method using the ESM-C protein language model. ESM-C achieved high accuracy for detecting variants of known targets (F1 scores >97%), with performance similar to BLASTP. However, ESM-C proved to be more effective at capturing functional relationships, distinguishing benign mutations from damaging ones where BLASTP did not. To characterize how screening would affect bioscience research, we measured flagging rates across diverse protein datasets. Flagging rates were significant for mammalian proteins weighted by publication frequency (23% for human, 20% for mouse), and rates for organisms distantly related to humans were minimal (<1.1% for bacteria, fungi, plants, and viruses). Among commercially relevant targets, 63% of antibody patent targets were classified as dual-use, reflecting that therapeutically important proteins often perform critical biological functions. To identify and flag risky user requests from protein binder design tools without placing an undue burden on scientific research and innovation, it will be essential to deploy this screening approach in a way that addresses the overlap our analysis showed between targets of concern and therapeutic targets-possibly in concert with tiered trusted access frameworks. This new method provides a foundation for proportionate safeguards for biological AI models that reduce misuse risks while preserving their benefits for legitimate research and demonstrates a concrete proof of principle that can be generalized to other protein design tools and biological AI models.

19
Protein size and geometry govern mutational robustness

Acharya, S.; Dey, S.

2026-08-20 bioinformatics 10.64898/2026.08.20.745934 medRxiv
Top 0.3%
9.3%
Show abstract

A fundamental question in structural biology centres around understanding protein evolution. Key to this process is mutational robustness, defined as the protein fold's ability to absorb sequence changes without collapsing its structure. Here, we show that robustness is systematically shaped by simple features such as protein size, geometry, and oligomeric state. We used Foldseek-identified (structural) homologs to quantify family size across monomers and higher homo-oligomers. We found that proteins in larger families are consistently larger in size, more compact in atomic density, and less exposed to solvent. Strikingly, homo-oligomers occupy systematically larger families than monomers, revealing quaternary structure itself as a driver of mutational tolerance, not merely a functional supplement. This signature of robustness can be further linked to increasing functional complexity in proteins; those with adaptive, multifaceted biological roles belong to larger structural families than those with specific roles, thereby linking structural flexibility directly to evolutionary versatility. In short, simple yet overlooked features of protein geometry can explain mutational robustness and evolvability, offering a structural rationale for why certain protein families have diversified extensively while others remain in evolutionary stasis.

20
In-Cell Protein Crystallization via a Locally Flexible 24-mer Assembly Precursor

Abe, S.; Tanaka, J.; Kikuchi, K.; Furuta, T.; Aizawa, Y.; Tanaka, Y.; Yokoyama, T.; Kanamaru, S.; Kobayashi, R.; Ueno, T.

2026-08-28 biophysics 10.64898/2026.08.25.746898 medRxiv
Top 0.4%
8.7%
Show abstract

In-cell protein crystallization (ICPC) produces ordered protein crystals within living cells, but the mechanisms used by proteins to acquire long-range crystalline order in the cellular environment remains poorly understood. Here, we define the assembly pathway of CipB, a crystalline inclusion protein from Photorhabdus luminescens. CipB crystals formed in cells dissolve under mild acidic conditions into a predominant 24-mer species, supporting a model in which an in-cell crystal is built from a discrete 24-mer assembly precursor rather than through direct packing of smaller oligomeric states. Structural analysis of recrystallized CipB shows that the same 24-mer architecture packs into a body-centered cubic lattice, consistent with the lattice observed for the in-cell crystals. Cryo-EM and molecular dynamics analyses indicate that the 24-mer assembly precursor preserves its overall architecture while retaining local conformational flexibility at the N-terminal and surface-loop regions. Mutation analyses further link the N-terminal region to the formation of the 24-mer precursor and surface residues to lattice assembly. These observations support a stepwise crystallization model in which N-terminal flexibility facilitates the formation of an assembly-competent 24-mer precursor, whereas defined hydrophobic surface contacts subsequently organize these precursors into a long-range-ordered lattice.